Why Even "95% Accurate" Agents Are Usually Wrong


The Arithmetic Nobody Runs

Welcome to this edition of QuipuCFO.

Think a 95% accuracy rate is a passing grade? In the world of Agentic AI, it’s a failure.

When you chain eight probabilistic steps together, that 95% compounds down to just 66% end-to-end. For every three tasks your agent processes, one is likely wrong. Independent benchmarks put real-world end-to-end accuracy closer to 11%.

This edition addresses this specific gap in financial reporting governance. I show how to overcome shortfalls in general IT controls in agentic AI.

Stop compounding errors from hitting your General Ledger by process decomposition and building "human-in-the-loop" controls.

Why Even "95% Accurate" Agents Are Usually Wrong

European insurers pay €2.8 billion a day in claims. At 95% per-step accuracy, an 8-step agentic workflow gets one in three wrong.


PwC and Microsoft published Unlocking Tomorrow, a playbook on agentic AI in financial services. It describes autonomous agents orchestrating claims journeys from first notice of loss to payment initiation, multi-agent platforms handling drawdowns, rollovers and repayments, and straight-through underwriting below defined thresholds.

The governance chapter covers regulations on Responsible AI, DORA, and the EU AI Act, and calls for a control tower to supervise agent behaviour. This edition is about challenges agentic AI poses to financial reporting and how to address those.

The arithmetic nobody runs

Every AI-agent demo comes with an accuracy number. The number is almost always a single-step figure: 85%, 95%, 99%. The number is real. It is also misleading.

Multi-step agentic workflows compound probability at every step. A workflow with eight steps running at 95% per-step accuracy produces end-to-end correctness of roughly 66%. Patronus AI's TRAIL benchmark puts real-world multi-step accuracy even lower: around 11% on realistic trace, implying 73% per-step accuracy results in 1 out of 9 workflows completing successfully end-to-end. The arithmetic illustrates how even small per-step error rates compound across Agentic workflows. CFOs who engage with the governance architecture now, before deployments reach financial reporting processes, are in a different position from those who engage after.

Why traditional controls don't reach

Deloitte's Q4 2025 CFO Signals survey of 200 finance chiefs at companies with at least $1 billion in revenue finds that 87% of CFOs believe AI will be extremely or very important to their finance department's operations in 2026, and 54% cite integrating AI agents as a top transformation priority. At the same time Deloitte reported only 1 in 5 organisations have a mature governance model in place for autonomous AI agents. The question is how control architecture in both upstream and finance processes will impact financial reporting.

Compounding error rates change the risk calculus in three ways that traditional IT General Controls do not cover.

First, the reliability chain breaks. Traditional validation depends on the principle that the same input produces the same output. Reconcile outputs to inputs, recalculate results, verify against source documents. Probabilistic systems break that principle by design. The control framework has to shift from "verify the logic once and rely on it" to "monitor the distribution of outputs and detect drift before it becomes material."

Second, errors cascade rather than isolate. A deterministic system fails in one place. A probabilistic agent operating in a chain can produce outputs that remain individually plausible but collectively wrong, with the error introduced at step three propagating through steps four through eight without any single step looking anomalous. Monitoring has to happen continuously throughout the process chain, not just at the point where outputs reach the ledger.

Third, most of these agents sit upstream of finance. The claims agent, the underwriting agent, the dynamic pricing agent. Finance does not build them or own them. Finance may not know when they are modified. But finance owns the misstatement risk when they fail. Zillow took a $304 million write-down in 2021 on an upstream valuation model that finance did not build. The control failure was not in the algorithm. It was in the absence of a control architecture connecting a probabilistic upstream system to the balance sheet it eventually touched.

The architecture that closes the gap

The control problem is structural, and it has a structural answer. The architecture has three components that work together.

Process decomposition before agent selection. Before evaluating any agentic solution, decompose the target process into discrete tasks. As discussed in the previous edition, for each task, determine whether it belongs in deterministic automation, probabilistic AI, or human judgment. The compounding error math applies only to the probabilistic steps. Finance workflows are almost always mixtures. The first governance move is knowing which steps are which, because the control framework differs for each. An invoice extraction step is deterministic. A coverage calculation step that draws on policy terms, claim history, and behavioural signals is probabilistic. Treating them identically in governance is the gap.

Governance intensity matched to agent profile. Not all agents require identical oversight. Classification should run across four dimensions: autonomy level, predictability of outputs, decision authority, and system integration scope. An invoice data extraction agent with constrained autonomy and predictable outputs warrants lighter governance than a treasury cash optimisation agent with high autonomy and multi-system reach. The classification determines what happens next: what approval gates are required, what escalation triggers are defined, and what override protocols allow finance to intervene. These elements belong in the agent's architecture from inception. Retrofitting governance onto deployed agents creates gaps and introduces friction at the worst possible time.

Dual-path monitoring that detects drift before it reaches the ledger. The PwC/Microsoft control tower concept is positioned as infrastructure for recording and monitoring agentic workflows. That framing is useful but incomplete for ICFR purposes. The monitoring architecture needs two distinct paths.

Hot path monitoring captures operational signals in near real-time: retrieval latency, errors, rejected operations, governance violations. It answers the question "is something failing right now?" Warm path monitoring aggregates telemetry over time to detect patterns that don't manifest as discrete failures: response consistency drift, reasoning loops, gradual efficiency decline, increasing exception rates. It answers the question "is performance degrading in a way that will matter at period close?"

The distinction matters for financial reporting because the failure mode most relevant to ICFR is not the dramatic one. It is the slow drift. A model that processes claims correctly 95% of the time in January, 92% in March, and 88% in June may never trigger a hot path alert. It will, eventually, produce a misstatement. Warm path monitoring is what catches that trajectory before the auditor does.

Where the human sits

PwC recommends a human-in-the-loop approach that requires human sign-off on agent actions exceeding defined thresholds based on risk factors, transaction values, or other considerations. That framing describes a threshold trigger. The governance question goes further: for which outputs does a human sit between the agent and a consequential decision, and for which outputs does a human supervise the aggregate rather than each transaction?

Human-in-the-loop is appropriate for high-materiality outputs where each individual decision carries misstatement risk: a large settlement determination, a treasury transaction above a defined threshold, a pricing adjustment to a material account. Human-on-the-loop is appropriate for high-volume outputs where the individual transaction is immaterial but the aggregate pattern is: claims frequency drift, coverage calculation distribution shifts, fraud flag rate changes. The control activity is different in each case. The first requires an approval gate. The second requires a monitoring protocol with defined escalation thresholds.

The question finance needs to answer for each agent deployment is not "is a human involved?" but "at what point does human accountability attach, and what information does that human have when it does?"

This week's articles

Agentic AI in Financial Services

Agentic AI is changing financial services by automating complex workflows, strengthening data governance and enabling secure, compliant AI adoption at scale. This playbook outlines the foundations needed to deploy it responsibly.
​

The measured leap Appraising AI agent impact with agent operations.

Explore approaches, metrics and AI agent observability solutions that can help your organization monitor and enhance AI agent performance against business and operational goals. By applying these recommendations, your organization can more rapidly progress on the path to agent-enabled workforce transformation—while protecting against unexpected costsand business risks.
​

Monday Moves: From Insight to Action

The investment signal in the PwC survey suggests these deployments are coming. The governance architecture question is whether finance engages with it now or explains the gap later.

Four questions to review your Agentic AI programme.

  1. Which of our agent deployments produce outputs that flow into the general ledger, directly or upstream? The EU AI Act inventory requirement is a starting point, but mapping inventory to material account balances is a separate exercise.
  2. For each deployment, which steps are genuinely probabilistic? This is the process decomposition question. Concentrate monitoring architecture where the compounding math applies.
  3. What detects drift in the probabilistic steps before it becomes material to a reported balance? Detection has to happen inside the period. Model monitoring, baseline comparisons, output distribution analysis. These control activities restore the reliability chain that probabilistic outputs break.
  4. If the agent is modified through retraining, redefining parameters, or prompt changes, does finance know within the close cycle? Traditional change management assumes discrete events. Agentic systems update continuously. The governance question is whether finance finds out from IT, from operations, or from the auditor.

Audit-proof your AI strategy.

Don't let your agentic deployments outpace your internal controls. Take the free QuipuCFO AI Readiness Assessment to see how your current governance with AI. You'll receive a personalised scorecard with actionable steps to bridge the gap between innovation and control. https://quipucfo.com/assessment]

​
​Unsubscribe · Preferences​